Search CORE

6 research outputs found

The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements made on several fronts over the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 67 new languages, including 30 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g. missing gender and macron information. We have also amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet

Proceedings - University of Groningen

University of Groningen

ARTS repository - University of Groningen

Dissertations of the University of Groningen

UniMorph 4.0: Universal Morphology

Author: Batsuren Khuyagbaatar
Bella Gábor
Budianskaya Elena
Coler Matt
Cotterell Ryan
El-Khaissi Charbel
et al.
Gasser Michael
Ghanggo Ate Yustinus
Goldman Omer
Gorman Kyle
Habash Nizar
Khalifa Salam
Kieraś Witold
Lane William Abbott
Leonard Brian
Mielke Sabrina
Montoya Samame Jaime Rafael
Nicolai Garrett
Pimentel Tiago
Raj Mohit
Ryskina Maria
Stöhr Niklas Werner
Publication venue: European Language Resources Association
Publication date: 01/01/2022
Field of study

Repository for Publications and Research Data

SIGMORPHON–UniMorph 2022 Shared Task 0: Generalization and Typologically Diverse Morphological Inflection

Author: Akkuş Faruk
Anastasopoulos Antonios
Andrushko Taras
Arora Aryaman
Atanalov Nona
Batsuren Khuyagbaatar
Bella Gábor
Budianskaya Elena
Cotterell Ryan
Dolatian Hossep
Ghanggo Ate Yustinus
Goldman Omer
Guriel David
Guriel Simon
Guriel-Agiashvili Silvia
Khalifa Salam
Kieraś Witold
Kodner Jordan
Krizhanovsky Andrew
Krizhanovsky Natalia
Marchenko Igor
Markowska Magdalena
Mashkovtseva Polina
Nepommiashchaya Maria
Rodionova Daria
Serova Alexandra
Sheifer Karina
Vylomova Ekaterina
Yemelina Anastasia
Young Jeremiah
Publication venue: Association for Computational Linguistics
Publication date: 01/07/2022
Field of study

The 2022 SIGMORPHON–UniMorph shared task on large scale morphological inflection generation included a wide range of typologically diverse languages: 33 languages from 11 top-level language families: Arabic (Modern Standard), Assamese, Braj, Chukchi, Eastern Armenian, Evenki, Georgian, Gothic, Gujarati, Hebrew, Hungarian, Itelmen, Karelian, Kazakh, Ket, Khalkha Mongolian, Kholosi, Korean, Lamahalot, Low German, Ludic, Magahi, Middle Low German, Old English, Old High German, Old Norse, Polish, Pomak, Slovak, Turkish, Upper Sorbian, Veps, and Xibe. We emphasize generalization along different dimensions this year by evaluating test items with unseen lemmas and unseen features separately under small and large training conditions. Across the five submitted systems and two baselines, the prediction of inflections with unseen features proved challenging, with average performance decreased substantially from last year. This was true even for languages for which the forms were in principle predictable, which suggests that further work is needed in designing systems that capture the various types of generalization required for the world’s languages

Repository for Publications and Research Data